Papers with Machine Translation
Copied to clipboard
| Challenge: | ACL 2019 received a record number of papers, close to three thousand, a sharp increase . aaron safina: organizers have to adapt quickly to numbers that surpass previous estimates . but he says conference needs to be large but should be enjoyable for all, retain original spirit . |
| Approach: | a record number of papers were submitted for the 2019 ACL conference . aaron ramirez: organizers worked hard to ensure a program that suits most participants . the conference will feature 9 tutorials, 18 one-day workshops and a varied program, he says . |
| Outcome: | a record number of papers received at the 2019 ACL, says cnn's nigel tsang . ttsong: organizers worked hard and professionally to make conference enjoyable for all . the conference will feature 9 tutorials, 18 one-day workshops, and the fourth conference on machine translation . |
Copied to clipboard
| Challenge: | . - (EN) |
| Approach: | . - (EN) |
| Outcome: | . - (EN) |
Copied to clipboard
| Challenge: | Despite improvements in machine translation output, remarkable differences can be observed when comparing machine translations (MT) and human translations. |
| Approach: | They describe a corpus of eye movement data collected during natural reading of a human translation and a machine translation of . they use this corpus to investigate the effect of machine translation on the reading process and the effects of various error types on reading. |
| Outcome: | The proposed corpus will be used in future research to investigate the effect of machine translation on the reading process and the effects of various error types on reading. |
Copied to clipboard
| Challenge: | a simple translation-test approach would fail the latency requirements of a live environment. |
| Approach: | They show that annotating unlabeled utterances offline can improve performance . they demonstrate that an extrinsic evaluation can improve the performance if manual data is available . |
| Outcome: | The proposed method improves performance in an extrinsic evaluation setting with real-world commercial dialog system in german. |
Copied to clipboard
| Challenge: | In recent years, the emergence of seq2seq models has revolutionized the field of machine translation by replacing traditional phrase-based approaches with neural machine translation (NMT) systems based on the encoder-decoder paradigm. |
| Approach: | They propose to use a convolutional seq2seq model to combine the strengths of the two approaches. |
| Outcome: | The proposed architectures outperform the existing models on the WMT’14 benchmark dataset. |
Copied to clipboard
| Challenge: | Recent advances in machine translation evaluations are expensive and time-intensive. |
| Approach: | They propose a method to evaluate the correlation between human and metric scores . they argue that it is equally important to ensure that metrics treat all systems fairly and consistently. |
| Outcome: | The proposed method ignores a central requirement of the evaluation process, and ignores the need for a thorough evaluation procedure. |
Copied to clipboard
| Challenge: | Existing studies have demonstrated that cross-lingual information retrieval performance is highly dependent on query translation quality. |
| Approach: | They investigate whether improving query translation quality yields little or no benefit to further improve retrieval performance. |
| Outcome: | The proposed methods compare query translations for multiple language pairs and identify the most promising language pairs to invest and improve. |
Copied to clipboard
| Challenge: | MT-Telescope is an open source, written in Python, and is built around a user friendly and dynamic web interface. |
| Approach: | They propose a platform to facilitate comparative analysis of the output quality of two Machine Translation (MT) systems. |
| Outcome: | The proposed platform supports fine-grained segment-level analysis and interactive visualisations that expose the fundamental differences in the performance of the compared systems. |
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) is a useful component in NLP applications. |
| Approach: | They propose to use annotated named entity corpora to classify a given entity into a category within a textual document. |
| Outcome: | The proposed model achieves an F1 score of 0.80 on an unseen dataset for Indian languages. |
Copied to clipboard
| Challenge: | Spoken Language Translation (SLT) evaluation of machine translation (MT) quality has been investigated for decades. |
| Approach: | They propose an open-source tool for assessing machine translation (MT) quality based on time-stamped transcripts and reference translations. |
| Outcome: | The proposed evaluation tool is based on time-stamped transcripts and reference translations into a target language. |
Copied to clipboard
| Challenge: | Existing LLMs are unable to match outputs due to their open-ended nature . |
| Approach: | They propose a meta-evaluation strategy PromptCUE to evaluate cutting-edge LAJ-MT models such as GEMBA-MQM and a rubric-style prompt tailored to the characteristics of LLMs. |
| Outcome: | The proposed model is able to predict scores or identify errors for individual sentences and is reliable in the real world. |
Copied to clipboard
| Challenge: | Existing approaches to recover dropped pronouns ignore the dependencies between pronounes in neighboring utterances. |
| Approach: | They propose a framework that combines Transformer network and General Conditional Random Fields to model the dependencies between pronouns in neighboring utterances. |
| Outcome: | The proposed framework outperforms state-of-the-art models on three Chinese conversation datasets showing that it captures the dependencies between pronouns in neighboring utterances. |
Copied to clipboard
| Challenge: | Sentence splitting is a major simplification operation. |
| Approach: | They propose a simple and efficient splitting algorithm based on an automatic semantic parser. |
| Outcome: | The proposed method compares favorably to the state-of-the-art in combined lexical and structural simplification. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated impressive few-shot learning capabilities through in-context learning. |
| Approach: | They propose a novel Alternating Minimization approach for example selection that improves ICL performance on low-resource Indic languages. |
| Outcome: | The proposed approach outperforms existing frameworks for retrieving examples on low-resource Indic languages. |
Copied to clipboard
| Challenge: | In recent years, there has been growing interest in voice-controlled devices, such as Amazon Alexa or Google home. |
| Approach: | They investigate the use of Machine Translation to bootstrap a natural language understanding system for a new language for the use case of a large-scale voice-controlled device. |
| Outcome: | The proposed method reduces the time and cost of getting annotated corpus for a new language while still providing a large enough coverage of user requests. |
Copied to clipboard
| Challenge: | LINGUIST generates annotated data for Intent Classification and Slot Tagging (IC+ST) we demonstrate fine-tuning of a large-scale seq2seq model to control outputs of multilingual data generation. |
| Approach: | They propose a method for generating annotated data for Intent Classification and Slot Tagging (IC+ST) they use a 5-billion-parameter multilingual sequence-to-sequence model to fine-tune it on a flexible instruction prompt. |
| Outcome: | The proposed method outperforms state-of-the-art approaches on a SNIPS intent setting and shows significant improvement on IC+ST in a cross-lingual setting. |
Copied to clipboard
| Challenge: | Unsupervised pre-training of large neural models has revolutionized Natural Language Processing. |
| Approach: | They propose to use pre-trained checkpoints for Sequence Generation to initialize a Transformer-based sequence-to-sequence model that is compatible with these checkpoint. |
| Outcome: | The proposed model is compatible with pre-trained BERT, GPT-2, and RoBERTa checkpoints and achieves state-of-the-art results on Machine Translation, Text Summarization, Sentence Splitting, and Sentance Fusion. |
Copied to clipboard
| Challenge: | Appraise is an open-source framework for crowd-based annotation tasks . it is used for shared tasks at the conference on machine translation and at IWSLT 2017 . |
| Approach: | They present an open-source framework for crowd-based annotation tasks . they describe the entire lifecycle of an Appraise evaluation campaign . |
| Outcome: | The proposed framework is used to run evaluation campaigns at the WMT Conference on Machine Translation and at IWSLT 2017 . it has been adopted by the translator team at Microsoft Translator for internal quality monitoring . |
Copied to clipboard
| Challenge: | Advancements in machine translation (MT) are hindered by the lack of dedicated parallel data, which are necessary to adapt MT systems to satisfy neutral constraints. |
| Approach: | They propose to use GPT-4 to generate GNTs that avoid bias and undue binary assumptions by comparing MT with the popular GPT-3 model. |
| Outcome: | The proposed model outperforms the existing model and provides valuable insights into the potential and challenges associated with prompting for neutrality. |
Copied to clipboard
| Challenge: | a number of Arabic text diacritizers use diacritics to convey information about meaning of a word . Arabic text to speech (TTS) requires a complex process to determine the correct diacritical for each character . |
| Approach: | They propose to use Arabic diacritization to enhance machine translation models . they propose to build automatic Arabic text diacritics using two approaches . |
| Outcome: | The proposed models are either better or on par with other models, which require language-dependent post-processing steps, unlike ours. |
Copied to clipboard
| Challenge: | Graph-based semantic parsing is one of the most promising general-purpose meaning representations . owing to this heterogeneity, most research focused on solutions specific to a given formalism . |
| Approach: | They propose a multilingual neural machine translation framework for Graph-based semantic parsing . they propose Graph2seq architecture that trains with an MNMT objective . |
| Outcome: | The proposed framework outperforms all competitors on cross-lingual parsing tasks. |
Copied to clipboard
| Challenge: | Existing approaches require large amounts of expert annotated data, computation, and time for training. |
| Approach: | They propose an unsupervised approach to QE where no training is required . they use a dataset that enables work on both black-box and glass-box approaches . |
| Outcome: | The proposed approach rivals state-of-the-art supervised QE models in terms of correlation with human judgments of quality. |
Copied to clipboard
| Challenge: | Neural Machine Translation models are sensitive to noise in the input data. |
| Approach: | They propose new methods to extend limited noisy data and further improve NMT robustness to noise while keeping the models small. |
| Outcome: | The proposed methods extend limited noisy data and improve robustness to noise while keeping the models small. |
Copied to clipboard
| Challenge: | Existing corpora of parallel corporata are being used in the biomedical domain . MT is known to support readers' access to textual documents in a language other than their native language . |
| Approach: | They propose to leverage parallel corpora to implement cross-lingual information retrieval or machine translation tools. |
| Outcome: | The proposed corpus is being used in the biomedical task at the conference on machine translation (WMT'16 and WMT'17) it can be leveraged to provide access to health information in languages other than English. |
Copied to clipboard
| Challenge: | IndicIRSuite is the first attempt at building large-scale Neural Information Retrieval resources for a large number of Indian languages. |
| Approach: | They introduce Neural Information Retrieval resources for 11 widely spoken Indian Languages from two major Indian language families. |
| Outcome: | Experiments show that Indic-ColBERT improves on INDIC-MARCO datasets for 11 languages, and that it can be used to improve IR for Indian languages. |
Copied to clipboard
| Challenge: | Noisy input text can cause disastrous mistranslations in most modern machine translation systems. |
| Approach: | They propose a benchmark dataset for Machine Translation of Noisy Text (MTNT) they use reddit comments and professionally sourced translations to examine noise types. |
| Outcome: | The proposed dataset can provide an attractive testbed for noise-robust machine translation systems. |
Copied to clipboard
| Challenge: | Existing models that capture speaker-related variations do not include explicit information about the speaker. |
| Approach: | They propose a method that adapts the bias of the output softmax to each particular user . they propose to model speaker-related variations as an additional bias vector in the softmax layer . |
| Outcome: | The proposed technique improves translation accuracy and better reflection of speaker traits in target text. |
Copied to clipboard
| Challenge: | a method to correct noisy User Generated Content (UGC) in French is proposed . it leverages on the existence of UGC specific noise due to the misuse of words with similar pronunciations. |
| Approach: | They propose a phonetizer-based method to correct noisy User Generated Content (UGC) they use phonetic similarity to generate IPA pronunciations of words . |
| Outcome: | The proposed method improves translation quality of noisy User Generated Content (UGC) in french. |
Copied to clipboard
| Challenge: | Existing approaches to training encoder-decoder systems often depend on teacher-forcing with the likelihood criteria, e.g. next token prediction of the reference sequence. |
| Approach: | They propose an inference-efficient way to modify the behaviour of an encoder-decoder system according to a specific attribute of interest by using a small proxy network. |
| Outcome: | The proposed framework improves the COMET performance of Flan-T5 on Machine Translation and the WER of Whisper foundation models on Speech Recognition. |
Copied to clipboard
| Challenge: | Most work on Knowledge Graph (KG) verbalisation is monolingual leaving open the question of how to scale KG-to-Text generation to languages with varying amounts of resources. |
| Approach: | They explore how to scale KG-to-Text generation to languages with varying resources . they construct multilingual training data and test data for each language . |
| Outcome: | The proposed approach performs best on all 9 languages, compared with other approaches on low vs high resource languages and on in- v. out-of-domain data. |
Copied to clipboard
| Challenge: | Recent work on fewshot prompting in Large Language Models has focused on selecting the few-shot samples for prompting. |
| Approach: | They propose a method to add demonstration attributes to prompting in machine translations by perturbations of high-quality in-domain demonstrations. |
| Outcome: | The proposed method improves upon the zero-shot translation performance of GPT-3, even making it competitive with few-shot prompted translations. |
Copied to clipboard
| Challenge: | In machine translation evaluation, metric performance is assessed based on agreement with human judgments. |
| Approach: | They incorporate human baselines into the MT meta-evaluation to gain a clearer understanding of metric performance and establish an upper bound. |
| Outcome: | The results suggest human parity, but there are several reasons to caution . |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have redefined Machine Translation, enabling context-aware and fluent translations across hundreds of languages and textual domains. |
| Approach: | They propose a framework and dataset to evaluate the translation quality and fairness of open-source LLMs. |
| Outcome: | The proposed framework and dataset evaluates translation quality and fairness of open-source LLMs. |
Copied to clipboard
| Challenge: | Potential Idiomatic Expression (PIE) dataset for NLP in English contains over 20,100 samples with almost 1,200 cases of idioms from 10 classes (or senses). |
| Approach: | They present a large Potential Idiomatic Expression (PIE) dataset for Natural Language Processing (NLP) in English. |
| Outcome: | The proposed dataset contains over 20,100 samples with almost 1,200 cases of idioms (with their meanings) from 10 classes (or senses). |
Copied to clipboard
| Challenge: | Nigeria is a multilingual country with 500+ languages. |
| Approach: | They propose to use a pidgin and a creole to analyze the pidgins of Nigeria . they also use machine translation to analyze their results . |
| Outcome: | The results show that the two pidgins do not represent each other and are hard to teach . the results show the pidgin varieties are underrepresented in Generative AI . |
Copied to clipboard
| Challenge: | Autoregressive decoding is expensive for many sequence-to-sequence tasks, but for some downstream tasks, the actual decoding output is not needed, just attributes of the sequence. |
| Approach: | They propose non-autoregressive proxy models that can efficiently predict scalar-valued sequence-level attributes from the encodings, avoiding the expensive decoding stage. |
| Outcome: | The proposed models outperform ensembles in machine translation (MT) and automatic speech recognition (ASR) while being significantly faster. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are general-purpose language models capable of many natural language generation or understanding tasks. |
| Approach: | They investigate how LLMs differ qualitatively from standard Neural Machine Translation models by measuring literalness and monotonicity. |
| Outcome: | The proposed models achieve close to state-of-the-art translation performance under few-shot prompting . the results are backed up by human evaluations and a newer MT quality metrics . |
Copied to clipboard
| Challenge: | Despite many recent successes, Machine Translation still lacks support or fails to achieve good performance for most low-resource languages. |
| Approach: | They propose a benchmark to evaluate OCR systems on low-resource languages and low- resource scripts. |
| Outcome: | The proposed benchmark evaluates state-of-the-art OCR systems on low-resource languages and low-rural scripts. |
Copied to clipboard
| Challenge: | Pronouns are often dropped in conversational genres as their referents can be easily understood from context. |
| Approach: | They propose an end-to-end neural network model to recover dropped pronouns in conversational data. |
| Outcome: | The proposed model improves on three different conversational genres. |
Copied to clipboard
| Challenge: | Using fine-grained evaluation techniques, translation outputs have become better and more fluent. |
| Approach: | They propose a fine-grained test suite for the language pair German–English . they describe the creation and implementation of the test suite in detail . |
| Outcome: | The proposed test suite is based on linguistically motivated categories and phenomena and semi-automatic evaluation is carried out with regular expressions. |
Copied to clipboard
| Challenge: | Existing methods to generate knowledge graphs are unable to handle non-English textual information. |
| Approach: | They propose a task of automatic Knowledge Graph Completion to bridge the gap between English and non-English textual information. |
| Outcome: | The proposed method bridges the gap between the quantity and quality of textual information between English and non-English languages. |
Copied to clipboard
| Challenge: | a new method for abstractive summarization is being developed for document summarizing . abstractive methods require extensive natural language generation to rewrite the sentences . |
| Approach: | They propose an unsupervised abstractive summarization system in multi-document context . they use a paraphrastic sentence fusion model which performs sentence synthesis and paraphrazing . |
| Outcome: | The proposed model improves information coverage and abstractiveness of generated sentences. |
Copied to clipboard
| Challenge: | Multi-way parallel, machine generated content dominates the translations in lower resource languages . a limited investigation suggests this selection bias is the result of low quality content generated in English and translated into many lower resource language via MT. |
| Approach: | They show that multi-way parallel, machine generated content dominates translations in many languages . they also find evidence of a selection bias in the type of content which is translated into many languages. |
| Outcome: | The results suggest that the low quality of multi-way translations on the web was likely created using machine translation. |
Copied to clipboard
| Challenge: | Existing models cannot identify exact metaphorical words within a sentence . current models do not rely on hand-crafted knowledge for training . |
| Approach: | They propose an unsupervised learning method that identifies and interprets metaphors at word-level without preprocessing. |
| Outcome: | The proposed method outperforms baseline models in two translation systems for English to Chinese showing that it paraphrases metaphors into their literal counterparts. |
Copied to clipboard
| Challenge: | Reliably evaluating Machine Translation (MT) through automated metrics is a long-standing problem. |
| Approach: | They propose to use MT models to generate multiple diverse translations and use them as surrogates to reference translations to obtain a quantification of translation variability. |
| Outcome: | The proposed approach improves correlation with human judgements of quality by 15%. |
Copied to clipboard
| Challenge: | Existing work has only explored textual context. |
| Approach: | They propose to use visual and text modalities to explore Quality Estimation for Machine Translation and integrate them into multimodal QE frameworks. |
| Outcome: | The proposed approaches improve on sentence-level and document-level predictions using visual features extracted from images. |
Copied to clipboard
| Challenge: | Existing APE and QE combination strategies have not shown significant performance gains in the field of automatic post-editing (APE). |
| Approach: | They propose to train a model on APE and QE tasks to improve the APE performance by using a multi-task learning methodology that treats both tasks as a 'bargaining game' they also investigate various existing combination strategies and show that their approach achieves state-of-the-art performance for a ‘distant’ language pair, viz., English-Marathi. |
| Outcome: | The proposed model improves on two different language pairs, viz., English-Marathi and English-German. |
Copied to clipboard
| Challenge: | Using the COVID-19 related dataset, we generated parallel corpora in 26 languages and 26 languages. |
| Approach: | They propose to exploit COVID-19 related metadata to generate parallel corpora using the EMM/MediSys processing chain of news articles. |
| Outcome: | The proposed corpora were based on the COVID-19 related dataset created with the Europe Media Monitor (EMM) / Medical Information System (MediSys) . |
Copied to clipboard
| Challenge: | a study of 14 Indian languages shows that cognates can be detected by word embeddings . cognates are variants of the same lexical form across languages . |
| Approach: | They propose to use cross-lingual word embeddings to detect cognates among 14 Indian languages . they then evaluate the impact of their method on neural machine translation . |
| Outcome: | The proposed method improves on a dataset of 12 Indian languages . it also improves quality of the extracted cognates by up to 2.76 BLEU . |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are introducing a new phase in machine translation . despite advances in MT, there are still many challenges to overcome . |
| Approach: | They propose to highlight several new directions for MT that are influenced by Large Language Models like GPT-4 and ChatGPT. |
| Outcome: | The proposed models offer vast linguistic understandings and bring innovative methodologies, such as prompt-based techniques, that have the potential to further elevate MT. |
Copied to clipboard
| Challenge: | Existing evaluation metrics for natural language generation are expensive and time-consuming. |
| Approach: | They propose a framework that utilizes LLMs to generate synthetic evaluation datasets . they propose meta-correlation to measure alignment between metric rankings and human benchmarks based on synthetic data . |
| Outcome: | The proposed framework achieves meta-correlations exceeding 0.9 in multilingual QA and replaces human judgment with synthetic evaluation datasets. |
Copied to clipboard
| Challenge: | Word-level Quality Estimation (QE) of Machine Translation aims to detect potential translation errors in the translated sentence without reference. |
| Approach: | They propose to use a human-generated translation judgment to generate a word-level quality estimate (QE) using a translation error rate toolkit to detect translation errors without reference. |
| Outcome: | The proposed dataset is more consistent with human judgment and confirms the effectiveness of the proposed tag-correcting strategies. |
Copied to clipboard
| Challenge: | Existing routing strategies rely on heuristics, external predictors, or absolute quality estimation to capture whether the large model provides a worthwhile improvement over the small one. |
| Approach: | They propose a budget allocation problem for routing large model to large model . they propose heuristics, external predictors, or absolute quality estimation to determine the optimal signal for budgeted decisions. |
| Outcome: | The proposed model outperforms heuristics, quality/difficulty estimation baselines and achieves a superior quality–budget Pareto frontier. |
Copied to clipboard
| Challenge: | Existing methods to remove sentences consisting of illegal characters are tedious and repetitive. |
| Approach: | They propose a statistical method to identify illegal characters in natural language processing . they use a fixed-size feature vector to generate a Gaussian mixture model for each sentence . |
| Outcome: | The proposed method can score sentences and filter corpus on clean corpus and improve performance. |
Copied to clipboard
| Challenge: | Existing parallel datasets limit multilingual evaluations due to the nature of linguistic annotation, which is tedious, subjective, and costly. |
| Approach: | They extend the Cross-lingual Natural Language Inference corpus with Croatian and use Facebook's 1.2B parameter m2m_100 model to analyze the train set and compare its quality with the existing machine-translated German set. |
| Outcome: | The proposed model is consistent with other XNLI dubs and is compared with the existing machine-translated German train set. |
Copied to clipboard
| Challenge: | Especially the trend towards neural MT has renewed peoples' interest in better and more analytical diagnostic methods for MT quality. |
| Approach: | They propose a framework that supports a linguistic evaluation of machine translations using test suites. |
| Outcome: | The proposed framework supports linguistic evaluation of (machine) translations using test suites. |
Copied to clipboard
| Challenge: | Existing methods to model and leverage semantic units in natural language do not provide a complete understanding of the whole sentence. |
| Approach: | They propose a method which models the integral meanings of semantic units within a sentence . they propose 'word pair encoder' to help identify the boundaries of semantic unit boundaries . |
| Outcome: | The proposed method outperforms baselines and supports the semantic unit representation of subwords and tokens. |
Copied to clipboard
| Challenge: | Neural encoder-decoder models have been successful at a variety of NLP tasks, including machine translation, parsing, and dialog generation. |
| Approach: | They propose a method for search-aware training via a continuous relaxation of beam search to enable global normalization. |
| Outcome: | The proposed approach is able to train globally normalized recurrent sequence models through simple backpropagation. |
Copied to clipboard
| Challenge: | Machine Translation (MT) systems based on fine-tuned large language models (LLMs) are at a higher risk of generating hallucinations, which can severely undermine user’s trust and safety. |
| Approach: | They propose a method that intrinsically learns to mitigate hallucinations during the model training phase. |
| Outcome: | The proposed method reduces hallucinations by 89% on an average across three unseen target languages while preserving translation quality. |
Copied to clipboard
| Challenge: | a new CLTL model is proposed to facilitate cross-linguistic transfer learning between distant languages . a key to CLTL is to learn a shared representation space for the given source-target language pair. |
| Approach: | They propose a new CLTL model that integrates machine translation with MT . they use an unannotated data technique to make use of the model's pre-training and fine-tuning . |
| Outcome: | The proposed model achieves better CLTL performance than the baseline model without more annotated data. |
Copied to clipboard
| Challenge: | Existing studies have shown that existing models amplify biases observed in training data. |
| Approach: | They propose to use MT and NLP to amplify biases observed in training data to investigate how bias amplification might affect language in a broader sense. |
| Outcome: | The proposed model amplifys biases observed in training data and could lead to an artificially impoverished language, the authors show. |
Copied to clipboard
| Challenge: | Recent work on MT robustness has demonstrated the need to build or adapt systems that are resilient to such noise. |
| Approach: | They propose to synthesize natural noise in social media data to enhance robustness of MT systems by leveraging natural noise. |
| Outcome: | The proposed method can make a vanilla MT system more resilient to noise, partially mitigating loss in accuracy resulting therefrom. |
Copied to clipboard
| Challenge: | Davidson et al.: hate speech is a growing media presence, but research on generating CNs has been limited . he says a new dataset for CN generation is available for basque and spanish . this dataset is based on a multilingual encoder-decoder model . |
| Approach: | They propose a new Basque and Spanish dataset for automatic CN generation . they use machine translation and professional post-edition to generate CNs in both languages . |
| Outcome: | The proposed datasets show that training on post-edited data improves generation over monolingual settings . similar results in zero-shot crosslingual evaluations show multilingual data augmentation outperforms training in English and Spanish . |
Copied to clipboard
| Challenge: | 'Low-resourced'-ness is a complex problem that goes beyond data availability and reflects systemic problems in society. |
| Approach: | They propose to use machine translation to scale to low-resourced languages by using a dataset and a benchmarking system to measure their resource use. |
| Outcome: | The proposed approach allows participants without formal training to make a unique scientific contribution. |
Copied to clipboard
| Challenge: | Recent success of Recurrent Neural Networks (RNNs) in Machine Translation (MT) has prompted attention mechanisms to be used in machine translation. |
| Approach: | They propose a tree-structured attention model on Tree Long Short-Term Memory Networks . they also experiment with three LSTM variants: bidirectional-LSTMs, Constituency Tree-LSTS, and Dependency Tree LSTS. |
| Outcome: | The proposed model is based on tree-LSTMs, constituency trees, and dependencies trees. |
Copied to clipboard
| Challenge: | Multiword Expressions (MWEs) are hard nuts for many natural language processing tasks. |
| Approach: | They annotate 28 types of Chinese MWEs and then examine 31 MTE metrics on groups of sentences containing different MWE. |
| Outcome: | The results show that MT systems and MTE metrics still suffer from MWEs . |
Copied to clipboard
| Challenge: | End-to-end Speech Translation (E2E ST) encoders lack global context representation, whereas MT encoder lacks it. |
| Approach: | They propose a Stacked Acoustic-and-Textual Encoding method for speech translation . they propose an adaptor module to alleviate representation inconsistency . |
| Outcome: | The proposed method achieves state-of-the-art BLEU scores of 18.3 and 25.2 on two ST tasks. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have advanced machine translation (MT) a meta-evaluation dataset focused on non-literal translations is lacking . experimental results show the inaccuracies of traditional MT metrics and the limitations of LLM-as-a-Judge. |
| Approach: | They propose a meta-evaluation framework that leverages sub-agents to evaluate machine translation metrics. |
| Outcome: | The proposed framework improves on the knowledge cutoff and score inconsistency problem. |
Copied to clipboard
| Challenge: | Existing approaches to stream ST combine advances in ASR and MT to achieve high quality translations without compromising the speed of the system. |
| Approach: | They propose to concatenate an Automatic Speech Recognition system followed by a Machine Translation system. |
| Outcome: | The proposed models improve on the Europarl-ST dataset on the BLEU score. |
Copied to clipboard
| Challenge: | Existing approaches to zero-shot learning with large language models are poorly calibrated for zero-shoot tasks. |
| Approach: | They propose a contrastive decoding objective with a decay factor to address in-context bias . they conduct experiments on 3 model types and sizes, 3 language directions, and beam search . |
| Outcome: | The proposed method outperforms state-of-the-art decoding objectives with 20 BLEU points improvement from the default objective in some settings. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have achieved impressive results in Machine Translation (MT). human evaluations reveal that LLM-generated translations still contain various errors. |
| Approach: | They propose a LLM-based self-refinement framework that feeds error information back into LLMs to facilitate self-finement, leading to enhanced translation quality. |
| Outcome: | The proposed framework outperforms internal refinement and feedback methods while ensuring a robust translation quality baseline. |
Copied to clipboard
| Challenge: | a new model for parallel sentence retrieval can be used to align parallel sentences in multilingual corpora . a faithful aligner can help narrow down the candidate pool without having to deal with an enormous search space . |
| Approach: | They propose a model that can be trained on only one language pair and transfers to low-resource languages with negligible degradation in performance. |
| Outcome: | The proposed model outperforms the previous model on the Tateoba dataset by 8.0 points in accuracy and using less than 0.6% of their parallel data. |
Copied to clipboard
| Challenge: | Existing machine translation models struggle with noisy data and tail-end words and phrases. |
| Approach: | They introduce and formalize a class of noise and variation that preserves meaning in the target language. |
| Outcome: | The proposed model can perform better on natural asemantic variation (NAV) the proposed model is robust to a variety of perturbations, but not all of them are achieved with organic variations. |
Copied to clipboard
| Challenge: | Existing methods to evaluate multiple systems are expensive and require human evaluators. |
| Approach: | They propose a novel online learning approach that dynamically converges to the top-3 ranked systems for the language pairs considered by taking advantage of human feedback. |
| Outcome: | The proposed approach converges to the top-3 ranked systems for the language pairs considered despite the lack of human feedback for many translations. |
Copied to clipboard
| Challenge: | supervised systems have not replaced dedicated supervised models for machine translation tasks. |
| Approach: | They propose to guide LLMs to post-edit MT with feedback from MQM annotations . they then fine-tune the LLM to improve its ability to exploit the feedback . |
| Outcome: | The proposed model improves TER, BLEU and COMET scores on Chinese-English, English-German and English-Russian data. |
Copied to clipboard
| Challenge: | State-of-the-art Neural Machine Translation models struggle with generating low-frequency tokens, tackling which remains a major challenge. |
| Approach: | They propose a loss function to better adapt model training to structural dependencies of conditional text generation by incorporating inductive biases of beam search into the training process. |
| Outcome: | The proposed method leads to significant gains over cross-entropy across different language pairs, especially on the generation of low-frequency words. |
Copied to clipboard
| Challenge: | Mixture-of-Experts (MoE) networks have been proposed as an efficient way to scale up model capacity and implement conditional computing. |
| Approach: | They propose a new architecture that combines multi-head attention with the MoE mechanism and a sparsely gated architecture that allows for faster computations. |
| Outcome: | The proposed architecture can scale up the number of attention heads and the number parameters while preserving computational efficiency. |
Copied to clipboard
| Challenge: | a recent study has found that Arabic is underrepresented in Large Language Models, especially in dialectal variations. |
| Approach: | They propose a benchmark for Arabic Dialect and Cultural Evaluation that evaluates Arabic dialect comprehension and generation. |
| Outcome: | The proposed model outperforms multilingual models on dialect comprehension and generation, but significant challenges persist in dialect identification, generation, and translation. |
Copied to clipboard
| Challenge: | Existing approaches to cognate detection use orthographic, phonetic and semantic similarity based features sets. |
| Approach: | They propose a method for enriching feature sets with cognitive features extracted from gaze behaviour data from human readers’ gaze behaviour. |
| Outcome: | The proposed method improves cognate detection performance by 10% and 12% over existing methods. |
Copied to clipboard
| Challenge: | Existing machine translation metrics give maximum score to reference sentence . however, these metrics overlook the possibility that candidate sentences outperform reference sentences in terms of quality. |
| Approach: | They propose a machine translation metrics that give an absolute score to a translated sentence based on the similarity with the reference sentence. |
| Outcome: | The proposed measure outperforms existing MT metrics in terms of quality and assigns positive scores to candidates that outperformed reference sentences. |
Copied to clipboard
| Challenge: | Lexical ambiguity poses one of the greatest challenges in the field of Machine Translation. |
| Approach: | They propose a new benchmark to study semantic biases in Machine Translation of nominal and verbal words in five different languages. |
| Outcome: | The proposed benchmark tests state-of-the-art machine translation systems against the new test bed and provides a statistical and linguistic analysis of the results. |
Copied to clipboard
| Challenge: | Recent advances in machine translation (MT) have improved performance on low-resource language pairs. |
| Approach: | They propose to freeze most BART parameters and add new ones to fine-tune a model trained on MT. |
| Outcome: | The proposed model outperforms naive fine-tuning on Vietnamese to English on a training set for Vietnamese to Vietnamese . the proposed model is able to fine- tune on smaller datasets while still maintaining the same model performance. |
Copied to clipboard
| Challenge: | Noisy texts are common in user-generated texts that appear abundant in social media platforms like SMS, online chat, email, blogs, wikis etc. |
| Approach: | They propose a sequence-to-sequence architecture that uses a gating mechanism to detect types of corrections required from English texts. |
| Outcome: | The proposed architecture performs better than non-gated models on machine translation and Summarization tasks. |
Copied to clipboard
| Challenge: | Multilingual neural machine translation models (MNMT) are effective on transferring knowledge between high-resource languages to low-resourced languages. |
| Approach: | They propose a multilingual multi-domain adapter which combines domain and language knowledge using meta-learning with adapters. |
| Outcome: | The proposed model outperforms other adapter methods in a domain shift and language pair translation task. |
Copied to clipboard
| Challenge: | afan oromo, amharic, and tigrinya are low-resourced languages . they are used for training, benchmarks, news, health, and sports . afono o'mara: quantity does not guarantee quality of MT datasets . |
| Approach: | They investigate the quality of machine translation datasets for three low-resourced languages . they found a large skew towards the male gender in the datasets . |
| Outcome: | The results show that training data has large representation of political and religious text, but benchmark datasets focus on news, health, and sports. |
Copied to clipboard
| Challenge: | Reinforcement Learning (RL) is dependent on the reward formulation due to the intrinsic difficulty of the task in the high-dimensional discrete action space and the sparseness of the standard reward functions. |
| Approach: | They propose a maximally dense semantic-level unsupervised reward function which mimics human evaluation by considering both sentence fluency and semantic similarity. |
| Outcome: | The proposed reward outperforms the standard sparse reward by 2% on average for in- and out-of-domain settings. |
Copied to clipboard
| Challenge: | a new study examines the degree of alignment between languages in multilingual embeddings . cross-lingual embeds are designed to encode linguistic concepts that bridge equivalent semantic meaning . a comprehensive approach is needed to address these questions. |
| Approach: | They employ clustering to uncover latent concepts within multilingual models . they introduce two metrics to quantify alignment and overlap of these concepts . |
| Outcome: | The proposed model can capture linguistic nuances across languages, but is not language-agnostic? the proposed model is able to capture nuances in multiple languages, the authors say. |
Copied to clipboard
| Challenge: | In this paper, we show that the combination of Phrase Pair Injection and Corpus Filtering boosts the performance of Neural Machine Translation systems. |
| Approach: | They propose to combine Phrase Pair Injection and Corpus Filtering to boost performance of Neural Machine Translation systems. |
| Outcome: | The proposed method improves machine translation models on low-resource language pairs . BLEU score improves over models trained with whole pseudo-parallel corpus augmented with parallel corpus. |
Copied to clipboard
| Challenge: | Existing bilingual or multi-lingual MWE corpora are limited for multilingual use . only 871 pairs of English-German MWEs are available for research . |
| Approach: | They present a collection of bilingual and multi-lingual MWEs extracted from parallel corpora. |
| Outcome: | The available bilingual or multi-lingual MWE corpus is very limited . the collection is a small collection of 871 pairs of English-German MWEs . |
Copied to clipboard
| Challenge: | Existing approaches to generalization to resource-rich languages are difficult . a recent study shows that word representations can be useful in low resource languages . |
| Approach: | They propose two approaches for improving generalization to low-resource languages by adapting continuous word representations using linguistically motivated subword units. |
| Outcome: | The proposed method improves generalization to low resource languages . it requires neither parallel corpora nor bilingual dictionaries and requires no parallel training . |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown significant potential as judges for Machine Translation (MT) quality assessment. |
| Approach: | They propose a framework that automatically post-edits the original translation based on each error, thereby filtering out non-impactful errors. |
| Outcome: | The proposed framework improves reliability and quality of error spans against GEMBA-MQM, across eight LLMs in both high- and low-resource languages. |
Copied to clipboard
| Challenge: | Existing methods for detecting hallucinations and omissions in Machine Translation systems focus on analyzing the model’s internal states or relying on external tools. |
| Approach: | They propose an Optimal Transport-based word aligner specifically designed to enhance the detection of hallucinations and omissions in Machine Translation systems. |
| Outcome: | The proposed method is competitive with state-of-the-art methods across 18 language pairs on the HalOmi benchmark and shows promising features. |
Copied to clipboard
| Challenge: | Cognates are words that have a common etymological origin and can facilitate the Second Language Acquisition (SLA) however, they also pose a challenge to various NLP applications such as Machine Translation and Cross-lingual Sense Disambiguation. |
| Approach: | They create two cognate datasets for twelve Indian languages and use them to generate cognate sets. |
| Outcome: | The proposed datasets are curated using previously available baseline cognate detection approaches and evaluated with the help of lexicographers. |
Copied to clipboard
| Challenge: | Linguistic Linked Open Data (LLOD) is a flourishing line of research in the language resource community . existing LLOD standards and vocabularies are not widely used in this community despite its popularity . |
| Approach: | They propose to use Linguistic Linked Open Data to link a Sumerian corpus with lexical resources . they use a linguistically annotated archive to create a corpus of cuneiform texts . |
| Outcome: | The proposed LLOD framework is used in assyriology, with philological resources underrepresented . the proposed framework is based on a linguistically annotated corpus of Sumerian texts . |
Copied to clipboard
| Challenge: | Public Multilingual Knowledge Management Infrastructure (PMKI) is a project launched by the European Commission to promote the digital single market in the EU. |
| Approach: | The paper presents the Public Multilingual Knowledge Management Infrastructure (PMKI) action launched by the European Commission to promote the Digital Single Market in the European Union. |
| Outcome: | The proposed public multilingual knowledge management infrastructure (PMKI) is a pilot project launched by the European Commission to promote the digital single market in the European Union. |
Copied to clipboard
| Challenge: | Recent work in cross-lingual learning has pivoted around multilingual models, which are typically pretrained on unlabeled corpora in multiple languages using some form of language modeling objective. |
| Approach: | They propose to use a stronger machine translation system to mitigat mismatch between training on original text and running inference on machine translated text. |
| Outcome: | The proposed approach is highly task dependent and calls into question the dominance of multilingual models for cross-lingual classification. |
Copied to clipboard
| Challenge: | Quality Estimation (QE) is an essential role in applications of Machine Translation (MT). |
| Approach: | They propose to fuse uncertainty quantification into a pre-trained cross-lingual language model to predict the translation quality. |
| Outcome: | The proposed method achieves state-of-the-art on the datasets of WMT 2020 QE shared task. |
Copied to clipboard
| Challenge: | Existing methods for seq2seq regularization use label smoothing, but it is difficult to extend it to other datasets. |
| Approach: | They propose a method that smooths over well formed relevant sequences that are semantically similar to the target sequence. |
| Outcome: | The proposed method shows a consistent and significant improvement over the state-of-the-art methods on different datasets. |
Copied to clipboard
| Challenge: | Qualitative assessments in the form of human judgements (question generation), attention visualization (MT), and sample output (summarization) provide further evidence of the ability of Scratchpad to generate fluent and expressive output. |
| Approach: | They propose to use the encoder as a "scratchpad" memory to keep track of what has been generated and guide future generation. |
| Outcome: | The proposed mechanism improves the fluency of seq2seq models on three well-studied natural language generation tasks. |
Copied to clipboard
| Challenge: | The study focuses on three LT areas of the greatest interest for the ECmachine translation (MT), speech technology, and cross-lingual search. |
| Approach: | This paper presents the key results of a competitiveness analysis of the European language technology market for three areas – Machine Translation, speech technology, and cross-lingual search. |
| Outcome: | The study focuses on three LT areas of the greatest interest for the ECmachine translation (MT), speech technology, and cross-lingual search. |
Copied to clipboard
| Challenge: | COMI-LINGUA is the largest manually annotated Hindi-English code-mixed dataset . 125K+ high-quality instances across five core NLP tasks are annotating by three bilingual annotators . |
| Approach: | COMI-LINGUA is the largest manually annotated Hindi-English code-mixed dataset . 125K+ high-quality instances are annotating by three bilingual annotators . |
| Outcome: | The dataset covers five core NLP tasks, including Token-level Language Identification, Matrix Language Identification and Named Entity Recognition. |
Copied to clipboard
| Challenge: | Existing approaches to integrate semantics into Natural Language Understanding (NLP) systems are cost-effective and environmental impact-related. |
| Approach: | They propose to provide semantically-annotated corpora for four NLU tasks across five languages and to drop the requirement of closed datasets. |
| Outcome: | The proposed model provides hundreds of millions of silver yet high-quality annotations for four NLU tasks across five languages. |
Copied to clipboard
| Challenge: | Recent work proposes neural machine translation (NMT) for Brazilian Portuguese. |
| Approach: | They propose a neural machine translation approach that generates equivalent sentences in target language and source language. |
| Outcome: | The proposed approach outperforms phrase-based statistical machine translation systems for some pairs of languages. |
Copied to clipboard
| Challenge: | a recent study shows that end-to-end systems are not structurally free. |
| Approach: | They propose to use dictionary- and word vector-based baselines to align NPs in the bitext . they argue that alignment of NP's in MT can be improved by using old-fashioned methods . |
| Outcome: | a new study shows that alignment of NPs in the bitext is relevant even in an end-to-end paradigm . the proposed system can be improved by bringing in old-fashioned methods, the authors argue . |
Copied to clipboard
| Challenge: | a demo demonstrates a system for quantitatively evaluating MT systems in isolation or multiple MT models collectively . performance of machine translation systems varies significantly with inputs of diverging features, such as genres, genres and surface properties. |
| Approach: | They propose a benchmarking interface that quantitatively evaluates MT systems in isolation or collectively . the interface can be extended to include additional filters such as lexical, morphological, and syntactic features. |
| Outcome: | The proposed system quantitatively evaluates MT systems on multiple domains and evaluation metrics. |
Copied to clipboard
| Challenge: | Existing corpora are used to train and evaluate machine translation systems, but little information is available about the methods used for producing the corpus, including translation direction. |
| Approach: | They used PubMed and publisher websites to obtain contact information for MEDLINE authors and asked about their abstract writing practices. |
| Outcome: | The authors of MEDLINE articles included in the English/Spanish, English/FR, and English/Portuguese (EN/PT) WMT 2019 test sets reported a response rate of over 20% . |
Copied to clipboard
| Challenge: | prevailing methods for machine translation are often hindered by misleading reward signals. |
| Approach: | They propose a framework that aligns large language models to human preferences . they propose 'M2PO' to correct the bias towards partial errors . |
| Outcome: | The proposed framework outperforms open-source models and achieves parity with proprietary models. |
Copied to clipboard
| Challenge: | linguistic diversity of India poses significant machine translation challenges, authors say . underrepresented tribal languages like Bhili lack high-quality linguistic resources . |
| Approach: | They introduce a Bhili-Hindi-English Parallel Corpus, the first and largest parallel corpus worldwide . they evaluated a wide range of proprietary and open-source MLLMs on bidirectional translation tasks . |
| Outcome: | The proposed corpus spans critical domains such as education, administration, and news. |
Copied to clipboard
| Challenge: | Speech-to-Text Translation systems rely on a sequential pipeline that combines ASR and MT models. |
| Approach: | They propose a parameter-efficient framework that integrates one LPSM with a multilingual MT engine. |
| Outcome: | The proposed framework integrates one LPSM with a multilingual MT engine. |
Copied to clipboard
| Challenge: | Existing datasets for machine translation quality estimation and post-editing have several shortcomings. |
| Approach: | They propose a dataset for machine translation quality estimation and automatic post-editing . they report the performance of baseline systems trained on the MLQE-PE dataset . |
| Outcome: | The proposed dataset contains human labels for up to 10,000 translations per language pair. |
Copied to clipboard
| Challenge: | a few million years ago, we lost our tails, and with them, the narratives and identities they carry. |
| Approach: | They argue that multilingual large language models are threatening our linguistic diversity . they argue that models collapse towards what is likely, driven by statistical biases . |
| Outcome: | a new paper shows that models can collapse and lead to loss of linguistic diversity . the authors argue that models are a field that values and protects multilingual diversity - a key challenge for the field . |
Copied to clipboard
| Challenge: | Neural Machine Translation (NMT) is a one-shot process that generates the target language equivalent of some source text from scratch. |
| Approach: | They propose a machine translation task which assumes an initial target sequence, that must be transformed into a valid translation of the source. |
| Outcome: | The proposed system outperforms other systems trained for similar tasks. |
Copied to clipboard
| Challenge: | Existing methods to generate semantic processors for languages lacking hand curated data are inefficiently slow and unaffordable in terms of human resources and economic costs. |
| Approach: | They propose to use statistical word alignments to project annotations from multiple sources to a target language. |
| Outcome: | The proposed method is effective to transport NER annotations across languages . it can generate a good statistical model for a new target language . |
Copied to clipboard
| Challenge: | Existing methods for detecting hallucinations in machine translation are limited for low-resource languages. |
| Approach: | They evaluate sentence-level hallucination detection approaches using Large Language Models (LLMs) they find that the choice of model is essential for performance. |
| Outcome: | The proposed models outperform the existing models in HRLs and LRLs on average by 0.16 MCC. |
Copied to clipboard
| Challenge: | Large-scale generative models can perform a wide range of NLP tasks using in-context learning. |
| Approach: | They aim to understand the properties of good in-context examples for machine translation in both in-domain and out-of-domain settings. |
| Outcome: | The proposed model outperforms a strong kNN-MT baseline in 2 out of 4 out-of-domain datasets. |
Copied to clipboard
| Challenge: | Recent efforts to improve the quality of machine-generated natural language content have been limited due to the large token usage required by complex evaluation prompts. |
| Approach: | They propose a prompt optimization approach that uses a smaller, fine-tuned language model to compress input data for evaluation prompt, thus reducing token usage and computational cost when using larger LLMs for downstream evaluation. |
| Outcome: | The proposed approach reduces token usage and costs by 2.37 compared with larger LLMs for downstream evaluation. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated impressive performance across numerous NLP tasks, but fine-tuning them for Machine Translation (MT) often introduces catastrophic forgetting, compromising the broad general abilities of LLMs and introducing potential security risks. |
| Approach: | They propose a method that harnesses the strong generative capabilities of Large Language Models to create rationales for training data, which are then "replayed" to prevent forgetting. |
| Outcome: | The proposed approach harnesses the strong generative capabilities of LLMs to create rationales for training data, which are then “replayed” to prevent forgetting. |
Copied to clipboard
| Challenge: | Existing evaluation methods focus on fluency and factual reliability, while neglecting figurative quality. |
| Approach: | They propose a set of human evaluation metrics focused on the translation of figurative language and a parallel metaphor corpus generated by post-editing. |
| Outcome: | The proposed evaluation protocol estimates four aspects of MT: Metaphorical Equivalence, Emotion, Authenticity, and Quality. |
Copied to clipboard
| Challenge: | Existing studies have shown promising results in multilingual translation with limited bilingual supervision. |
| Approach: | They propose a Language-Aware Neuron Detecting and Routing framework that fine tunes LLMs to Machine Translation with diverse translation training data. |
| Outcome: | The proposed framework selectively finetunes LLMs to MT tasks with diverse translation training data. |
Copied to clipboard
| Challenge: | a genetic algorithm generates adversarial translations for arbitrary metrics . produced translations score well in an arbitrary MT evaluation metric, despite serious, deliberately introduced errors. |
| Approach: | They propose a method for decoding translation candidates from a machine translation model via a genetic algorithm to generate adversarial translations to test and challenge MT evaluation metrics. |
| Outcome: | The proposed method scores very well in an arbitrary MT evaluation metric despite serious, deliberately introduced errors. |
Copied to clipboard
| Challenge: | Autoregressive decoding limits the efficiency of transformers for Machine Translation (MT) Existing methods to solve this problem are expensive and require changes to the model. |
| Approach: | They propose to reframe autoregressive decoding with a parallel formulation . they propose to speed up existing models without training or modifications while retaining translation quality. |
| Outcome: | The proposed model speeds up existing models without training or modifications while retaining translation quality. |
Copied to clipboard
| Challenge: | reproducibility of experiments is a key issue in Neural Networks, which are fed with variable samples of training data. |
| Approach: | They reproduce some of the experiments related to neural network training for Machine Translation as reported in . they annotated a sample from the EN-FR and EN-DE Europarl with syntactic and semantic annotations to train neural networks with the Nematus Neural Machine Translation toolkit. |
| Outcome: | The results obtained were lower than the original paper, but on a more limited set of annotations. |
Copied to clipboard
| Challenge: | Multilingual demands and accessibility have made MT a global tool . however, the understanding of MT consumed by such a diverse group of users remains limited. |
| Approach: | They first trace the evolution of MT user profiles, focusing on non-experts and how their engagement with technology may shift with the rise of LLMs. |
| Outcome: | The proposed approach will help to align MT with user needs and improve the quality of the language. |
Copied to clipboard
| Challenge: | a major challenge in the practical use of Machine Translation (MT) is that users lack information on translation quality to make informed decisions about how to rely on outputs. |
| Approach: | They evaluate quality estimation feedback in vivo with a human study in a medical setting. |
| Outcome: | The proposed method improves appropriate reliance on MT, but backtranslation helps detect harmful errors. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have been successful in machine translation, but lack of high-quality parallel corpora and cost constrain scalability. |
| Approach: | They propose an LLM-driven dual-learning framework that enables autonomous translation . they employ a robust semantic-aware reward function that balances adequacy with reconstruction fidelity . |
| Outcome: | The proposed model outperforms larger models on benchmarks and achieves parity with state-of-the-art supervised baselines on mainstream benchmarks. |
Copied to clipboard
| Challenge: | Existing models for low-resource languages often focus on creating the largest possible dataset for generic translation. |
| Approach: | They develop a dataset for the specific domain of health for a low-resource English to Irish language pair and compare it to other similar datasets. |
| Outcome: | The proposed model improved BLEU score by 22.2 points compared with top performing models from the LoResMT2021 Shared Task. |
Copied to clipboard
| Challenge: | We introduce two language models with 1.2 billion and 3.7 billion parameters to improve Machine Translation (MT) for low-resource languages. |
| Approach: | They propose a set of tools to improve Machine Translation (MT) for low-resource languages with a focus on African languages. |
| Outcome: | The proposed model outperforms existing models on MT for African languages and improves translation evaluation metrics for 1K languages including African languages. |
Copied to clipboard
| Challenge: | Existing studies use parallel corpora for training, which results in less diverse paraphrases. |
| Approach: | They train a bidirectional multilingual neural machine translation model on a bilingual parallel corpus and use it as a paraphrasing model. |
| Outcome: | The proposed method generates paraphrases with higher semantic consistency, literal fluency and sentential diversity than existing parabanks and LLMs. |
Copied to clipboard
| Challenge: | Neural Machine Translation models still require translation post-editing to rectify errors and enhance quality under critical settings. |
| Approach: | They use GPT-4 to automatically post-edit NMT outputs across several language pairs . they show that GPT4 is adept at translation post- editing, producing meaningful edits . |
| Outcome: | The proposed translation post-editor improves on state-of-the-art language models on English-Chinese, English-German, Chinese-English and German-English language pairs. |
Copied to clipboard
| Challenge: | a recent data sampling method skews the annotated data toward shorter documents, not necessarily representative of the full test set. |
| Approach: | They examine different approaches to human evaluation and ranking of machine translation systems at the conference on machine translation . they propose a method that uses available labour budget to sample data in a more representative manner . |
| Outcome: | The proposed method improves representation of document lengths and produces stable rankings of translation quality. |
Copied to clipboard
| Challenge: | Existing research on disfluency correction has primarily focused on English due to the unavailability of large-scale open-source datasets. |
| Approach: | They propose to use an annotated human-annotated corpus to analyze disfluency correction in four important Indo-European languages to demonstrate the benefits. |
| Outcome: | The proposed model improves BLEU scores by 5.65 points when used with a state-of-the-art machine translation system. |
Copied to clipboard
| Challenge: | Quality Estimation (QE) is the task of evaluating the quality of a translation when reference translation is unavailable. |
| Approach: | They propose a Quality Estimation based Filtering approach to extract high-quality parallel data from the pseudo-parallel corpus. |
| Outcome: | The proposed approach improves the machine translation system performance by up to 1.8 BLEU points over the baseline model. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated excellent performance in Machine Translation (MT) however, a performance gap persists between high-resource and low-resourced languages due to imbalanced pre-training data. |
| Approach: | They propose a layer-wise metric to quantify the activation divergence between high- and low-resource languages. |
| Outcome: | The proposed model outperforms standard LoRA fine-tuning on Chinese-to-seven low-resource language translations. |
Copied to clipboard
| Challenge: | Existing studies that analyze unseen domains vary translation systems, annotators, or evaluation conditions, confounding domain effects with human annotation noise. |
| Approach: | They propose to use human error span annotations to evaluate translations of six translation systems across one seen news domain and two unseen technical domains to address these biases. |
| Outcome: | The proposed model improves on the human annotations in two unseen domains and on the news domains. |
Copied to clipboard
| Challenge: | Recent studies have shown that MT metrics return assessments as scalar scores that are difficult to interpret, posing a challenge to making informed design choices. |
| Approach: | They propose an interpretable evaluation framework that evaluates MT metrics in two scenarios that serve as proxies for filtering and translation re-ranking use cases. |
| Outcome: | The proposed framework offers clearer insights than correlation with human judgments. |
Copied to clipboard
| Challenge: | Despite progress in MT, a gap persists between how the technology is developed and how it is used in real-world contexts. |
| Approach: | They propose a human-centered approach to machine translation (MT) they argue that MT should be evaluated with diverse goals and contexts of use . |
| Outcome: | The proposed approach emphasizes alignment of evaluation and design with diverse communicative goals and contexts of use. |
Copied to clipboard
| Challenge: | Existing metrics have been developed and validated for English and other languages . this narrow focus leaves Indian languages largely overlooked, casting doubt on universality of current evaluation practices. |
| Approach: | They propose a large-scale benchmark that compares 26 automatic metrics with human judgments across six major Indian languages. |
| Outcome: | ITEM evaluates alignment of 26 automatic metrics with human judgments across six languages . authors: outliers exert significant impact on metric-human agreement, improve fidelity . they say the results offer critical guidance for advancing metric design and evaluation in Indian languages - a global market for machine translation and text summarization systems. |
Copied to clipboard
| Challenge: | generative large language models (LLMs) can perform in-context learning . machine translation (MT) has been shown to benefit from in-constitu examples . |
| Approach: | They propose a compositional translation paradigm that replaces naive few-shot MT with similarity-based demonstrations. |
| Outcome: | The proposed paradigm replaces naive few-shot MT with similarity-based demonstrations. |
Copied to clipboard
| Challenge: | Existing methods to calculate agreement are acceq and acc eq, but there are many ways to compute agreement based on how scores are grouped together. |
| Approach: | They propose a segment-level meta-evaluation metric that utilizes pairwise differences rather than raw scores to refine Global Pearson to intra-segment comparisons. |
| Outcome: | The proposed metric correctly ranks sentinel evaluation metrics and better aligns with human error weightings than acceq. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated impressive results in Machine Translation by following instructions, even without training on parallel data. |
| Approach: | They propose a Translate After LEarNing Textbook approach which aims to enhance LLMs’ ability to translate low-resource languages by learning from a textbook. |
| Outcome: | The proposed approach improves translation performance by 14.8% using 112 low-resource languages from FLORES-200 with two LLMs: ChatGPT and BLOOMZ. |
Copied to clipboard
| Challenge: | Existing studies on large language models focus on literal-level translation quality, such as adequacy and fluency. |
| Approach: | They propose a Culture-Aware Novel-Driven Parallel Dataset for Machine Translation and a multi-dimensional evaluation framework for assessing cultural translation quality. |
| Outcome: | The proposed model improves evaluation reliability in LLM-as-a-judge scenarios under culture-aware constraints. |
Copied to clipboard
| Challenge: | ailsntua researchers examine whether machine translation systems exhibit gender biases that reinforce societal stereotypes. |
| Approach: | They propose a probability-based metric to evaluate gender bias by analyzing aggregated model responses. |
| Outcome: | The proposed metric evaluates whether translations in Greek and French align with or diverge from societal stereotypes. |
Copied to clipboard
| Challenge: | Using machine translation tools for everyday tasks is becoming more commonplace, but a lack of evaluation strategies and alternatives can cause users to over-rely on it. |
| Approach: | They propose to use MT evaluation techniques to promote MT quality and MT literacy among its users. |
| Outcome: | The findings highlight the need for evaluation and NLP explanation techniques to promote MT quality and MT literacy among its users. |
Copied to clipboard
| Challenge: | Large Language Models suffer from paraphrasing errors, omissions, or hallucinations when input contains translation-specific elements that require strict preservation or controlled transformation. |
| Approach: | They propose a Controllable Element-Oriented Machine Translation framework that decomposes the translation process into a linguistically grounded analysis, strategy formulation, and final generation. |
| Outcome: | The proposed framework improves on the WMT23/24 Chinese–English benchmarks while significantly reducing element-level constraint violations. |